Papers with evaluation frameworks

35 papers
Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: a large number of studies rely on closed-style multiple-choice surveys to evaluate cultural alignment in Large Language Models . however, these methods are constrained and lack nuanced and accurate evaluations based on specific cultural proxies.
Approach: They propose to use the World Values Survey and Hofstede Cultural Dimensions as case studies to examine cultural alignment in Large Language Models.
Outcome: The findings advocate for more robust evaluation frameworks that focus on cultural proxies.
Guardrails and Security for LLMs: Safe, Secure and Controllable Steering of LLM Applications (2025.acl-tutorials)

Copied to clipboard

Challenge: Pretrained generative models provide novel ways for users to interact with computers.
Approach: This tutorial provides an overview of key guardrail mechanisms developed for LLMs along with evaluation methodologies and a detailed security assessment protocol.
Outcome: This tutorial provides an overview of key guardrail mechanisms developed for LLMs, along with evaluation methodologies and a detailed security assessment protocol.
SNaC: Coherence Error Detection for Narrative Summarization (2022.emnlp-main)

Copied to clipboard

Challenge: SNaC framework is used to evaluate long summaries, but it fails to identify gaps in coherence . nallapati and colleagues have developed a framework for fine-grained annotations of long summarizations .
Approach: They propose a narrative coherence evaluation framework for fine-grained annotations of long summaries that can be used to evaluate coherent narratives.
Outcome: The proposed framework can support future work in document summarization and coherence evaluation, the authors show .
GlotEval: A Test Suite for Massively Multilingual Evaluation of Large Language Models (2025.emnlp-demos)

Copied to clipboard

Challenge: Existing evaluation frameworks focus on English and a handful of high-resource languages, thereby overlooking the realistic performance of large language models in multilingual and lower-resourced scenarios.
Approach: They propose a unified and lightweight framework that integrates 27 benchmarks under a standard ISO 639-3 language identifier system to enable seamless incorporation of new benchmarks.
Outcome: The proposed framework integrates 27 benchmarks under a standard ISO 639-3 language identifier system, allowing for seamless incorporation of new benchmarks.
Benchmark Transparency: Measuring the Impact of Data on Evaluation (2024.naacl-long)

Copied to clipboard

Challenge: In this paper, we quantify the impact that data distribution has on the performance and evaluation of NLP models.
Approach: They propose to use disproportional stratified sampling to measure the data distribution across 6 different dimensions to quantify model performance.
Outcome: The proposed framework measures the data distribution across 6 different dimensions and shows that it is statistically significant and predicts model performance.
Too Open for Opinion? Embracing Open-Endedness in Large Language Models for Social Simulation (2026.eacl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly used to simulate public opinion and other social phenomena.
Approach: They argue that open-endedness is essential for realistic social simulations . they argue that it captures expressiveness and individuality .
Outcome: The proposed frameworks can improve measurement and design, support exploration of unanticipated views, and reduce researcher-imposed directive bias.
ScEdit: Script-based Assessment of Knowledge Editing (2025.findings-acl)

Copied to clipboard

Challenge: Knowledge Editing (KE) has gained increasing attention, yet current evaluation frameworks do not integrate KE into real-world application scenarios.
Approach: They propose a script-based benchmark which encompasses both counterfactual and temporal edits and integrates token-level and text-level evaluation methods.
Outcome: The proposed method combines token-level and text-level evaluation methods with a new fact-based evaluation framework.
The Model Agreed, But Didn’t Learn: Diagnosing Surface Compliance in Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models internalize vast world knowledge as parametric memory, yet inherit the staleness and errors of their source corpora.
Approach: They propose a framework that subjects models to discriminative self-assessment under diverse contextual pressures to scrutinize subtle behavioral nuances induced by memory modifications.
Outcome: The proposed framework achieves high benchmarks without overwriting internal beliefs, while recursive modifications accumulate representational residues, triggering cognitive instability and permanently diminishing the reversibility of the model’s memory state.
Examining the State-of-the-Art in News Timeline Summarization (2020.acl-main)

Copied to clipboard

Challenge: Existing work on news timeline summarization (TLS) has left an unclear picture of how well it is currently solved and how it can be approached.
Approach: They propose a combination of different TLS strategies that improves over the stateof-the-art on all tested benchmarks.
Outcome: The proposed method improves over the state-of-the-art on all tested benchmarks.
HateXScore: A Metric Suite for Evaluating Reasoning Quality in Hate Speech Explanations (2026.eacl-long)

Copied to clipboard

Challenge: Existing evaluation frameworks do not assess why a text is deemed hateful . authors present a new metric to evaluate the reasoning quality of model explanations .
Approach: They propose a metric suite to evaluate the reasoning quality of model explanations.
Outcome: The proposed metric validates it as a practical tool for trustworthy and transparent moderation on six diverse hate speech datasets.
CDT: A Comprehensive Capability Framework for Large Language Models Across Cognition, Domain, and Task (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing benchmarks focus on isolated abilities, lacking a holistic framework for assessing LLM capabilities.
Approach: They propose a Cognition-Domain-Task framework which measures a model’s capabilities across three dimensions.
Outcome: The proposed framework improves performance on dataset evaluation and data selection, while achieving higher scores on general and specific benchmarks.
Learning from Impairment: Leveraging Insights from Clinical Linguistics in Language Modelling Research (2025.coling-main)

Copied to clipboard

Challenge: Using neurolinguistics and aphasiology, we examine the theoretical underpinnings of some influential linguistically motivated training approaches targeting the syntactic domain.
Approach: They examine the theoretical underpinnings of linguistically motivated training approaches derived from neurolinguistics and aphasiology to develop human-like learning strategies for language models.
Outcome: The proposed frameworks can be used to improve the recovery and generalization of linguistic skills in aphasia treatment and to develop human-like learning strategies.
FINEST: Improving LLM Responses to Sensitive Topics Through Fine-Grained Evaluation (2026.findings-eacl)

Copied to clipboard

Challenge: Existing evaluation frameworks lack systematic methods to identify weaknesses in LLMs . Existing methods to evaluate LLM responses to sensitive topics are lacking .
Approach: They propose a FINE-grained response evaluation taxonomy for sensitive topics that breaks down helpfulness and harmlessness into errors across three main categories: Content, Logic, and Appropriateness.
Outcome: The proposed model outperforms refinement without guidance on Korean-sensitive questions . FINEST significantly improves the model responses across all three categories .
BiRRE: Learning Bidirectional Residual Relation Embeddings for Supervised Hypernymy Detection (2020.acl-main)

Copied to clipboard

Challenge: supervised hypernymy detection has been studied under various frameworks . supervised classifiers are more likely to suffer from "lexical memorization"
Approach: They propose a representation learning framework called Bidirectional Residual Relation Embeddings to model the possibility of a term being mapped to another in the embedding space by hypernymy relations.
Outcome: The proposed model outperforms baselines over evaluation frameworks.
RAFFLES: Reasoning-based Attribution of Faults for LLM Systems (2026.eacl-long)

Copied to clipboard

Challenge: Existing evaluation frameworks focus on simple metrics and end-to-end outcomes, but they struggle with longer contexts.
Approach: They propose an offline evaluation architecture that incorporates iterative reasoning to evaluate the quality of the candidate faults and rationales of the Judge.
Outcome: The proposed architecture outperforms baseline evaluation frameworks with two datasets to identify step-level faults in multi-agent systems and ReasonEval datasets.
ReportLogic: Evaluating Logical Quality in Deep Research Reports (2026.acl-long)

Copied to clipboard

Challenge: Existing evaluation frameworks that evaluate large language models for Deep Research largely ignore this requirement.
Approach: They propose a benchmark that quantifies report-level logical quality through a reader-centric lens of auditability.
Outcome: The proposed model quantifies logical quality through a reader-centric lens of auditability.
Unanswerability Evaluation for Retrieval Augmented Generation (2025.acl-long)

Copied to clipboard

Challenge: Existing evaluation frameworks for retrieval-augmented generation (RAG) systems focus on answerable queries, but ignore the importance of appropriately rejecting unanswerable requests.
Approach: They propose a framework to evaluate whether retrieval-augmented generation systems handle unanswerable queries specific to a given knowledge base.
Outcome: The proposed framework synthesizes diverse and challenging queries for any given knowledge base and evaluates them with unanswered ratio and acceptable ratio metrics.
Multi-Dimensional Evaluation of Text Summarization with In-Context Learning (2023.findings-acl)

Copied to clipboard

Challenge: In-context learning-based evaluators are competitive with learned evaluation frameworks for text summarization tasks.
Approach: They propose to use large language models as multi-dimensional evaluators using in-context learning to evaluate text summarization tasks.
Outcome: The proposed frameworks are competitive with existing frameworks on relevance and factual consistency, the authors show .
A Survey of LLM-based Agents in Medicine: How far are we from Baymax? (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are transforming healthcare through their ability to understand and assist with medical tasks.
Approach: They analyze system profiles, clinical planning, medical reasoning frameworks, and external capacity enhancement.
Outcome: The findings highlight the future directions in medical reasoning, physical system integration, and training simulations.
LiveFact: A Dynamic, Time-Aware Benchmark for LLM-Driven Fake News Detection (2026.acl-long)

Copied to clipboard

Challenge: Current evaluation frameworks are static and vulnerable to benchmark data contamination . current models are ineffective at assessing reasoning under temporal uncertainty .
Approach: They propose a live-based benchmark that simulates the real-world "fog of war" they propose evaluating models on their ability to reason with evolving, incomplete information .
Outcome: The proposed model outperforms proprietary state-of-the-art models in classification and evidence mode . it also provides a component to monitor BDC explicitly .
Listen, Watch, and Learn to Feel: Retrieval-Augmented Emotion Reasoning for Compound Emotion Generation (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods to assess human emotion are limited by the subjective nature of emotion perception, limiting the robustness of existing models.
Approach: They propose a plug-and-play module that enhances MLLMs’ ability to tackle compound and context-rich emotion tasks.
Outcome: The proposed framework improves MLLMs' ability to tackle compound and context-rich emotion tasks and the Compound Emotion QA dataset shows it performs well across both benchmarks and evaluation frameworks.
Stereotype Bias in a Bilingual Setting: A Culturally Grounded Evaluation in Kazakhstan (2026.acl-long)

Copied to clipboard

Challenge: Stereotype bias in language models is largely understudied in English . language models perform strongly on downstream NLP tasks, but they are pre-trained on large text corpora .
Approach: They use a dataset to assess stereotype bias in language models in Kazakhstan . they find that stereotype bias is most pronounced in code-mixed inputs .
Outcome: The proposed dataset shows that stereotype bias is most pronounced in code-mixed inputs.
ICLEval: Evaluating In-Context Learning Ability of Large Language Models (2025.coling-main)

Copied to clipboard

Challenge: Existing evaluation frameworks focus on language abilities and knowledge, often overlooking the assessment of ICL ability.
Approach: They propose to evaluate the ICL ability of Large Language Models (LLMs) using the ICLEval benchmark.
Outcome: The proposed benchmark demonstrates that ICL ability is universally present in different LLMs and model size is not the sole determinant of ICL efficacy.
TextEE: Benchmark, Reevaluation, Reflections, and Future Challenges in Event Extraction (2024.findings-acl)

Copied to clipboard

Challenge: Recent studies suggest that event extraction evaluations may not accurately reflect the true performance.
Approach: They propose a standardized, fair, and reproducible benchmark for event extraction . they use standardized scripts and splits for 16 datasets spanning eight domains .
Outcome: The proposed benchmarks show that they struggle to achieve satisfactory performance.
OMHBench: Benchmarking Balanced and Grounded Omni-Modal Multi-Hop Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Existing evaluation frameworks for multimodal large language models suffer from limitations . modality shortcuts and biased reasoning paths are common in such models .
Approach: a new benchmark evaluates omni-modal multi-hop reasoning using 6,144 questions . authors propose OMHBench to address these limitations by comparing modalities .
Outcome: OMHBench evaluates omni-modal multi-hop reasoning on 6,144 questions with balanced reasoning paths . evaluation of 13 state-of-the-art models shows performance gap exists between MLLMs and open-source models .
Beyond the Last Frame: Process-aware Evaluation for Generative Video Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Existing evaluation frameworks often rely on single-frame assessments, which can lead to outcome-hacking.
Approach: They propose a process-aware evaluation paradigm that uses a hierarchical rubric to evaluate the validity of the intermediate steps and the final result.
Outcome: The proposed model achieves POC@1.0 only about 20% and exhibits significant outcome-hacking.
Financial Language Model Evaluation (FLaME) (2025.findings-acl)

Copied to clipboard

Challenge: Language Models (LMs) have demonstrated impressive capabilities with core NLP tasks in finance, but their effectiveness is difficult to assess due to gaps in evaluation methodologies.
Approach: They propose to use a framework to evaluate language models against ‘reasoning-reinforced’ LMs to measure their performance on finance NLP tasks.
Outcome: The proposed frameworks are open-source and provide data and data for the study.
Help Me Write a Story: Evaluating LLMs’ Ability to Generate Writing Feedback (2025.acl-long)

Copied to clipboard

Challenge: Current models provide specific and mostly accurate writing feedback, but they fail to identify the biggest writing issue in the story and to correctly decide when to offer critical vs. positive feedback.
Approach: They propose a task that corrupts 1,300 stories to intentionally introduce writing issues to study model performance.
Outcome: The proposed model performs well in a controlled task with human and automatic evaluation metrics.
Reranking-based Generation for Unbiased Perspective Summarization (2025.findings-acl)

Copied to clipboard

Challenge: Existing evaluation frameworks rely on traditional metrics for measuring key attributes such as coverage and faithfulness without verifying their applicability.
Approach: They propose to use human annotations to measure perspective summary quality and reranking-based methods yield strong results.
Outcome: The proposed methods show that they perform well with synthetically generated and reranking-labeled data.
MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing evaluation frameworks assess isolated responses using coarse-grained taxonomies or static datasets.
Approach: They propose a role-aware mental health safety taxonomy that characterizes clinically significant harm in terms of interactional roles an AI counselor adopts.
Outcome: The proposed framework significantly improves failure-mode coverage and diagnostic granularity.
PrefIx: Understand and Adapt to User Preference in Human-Agent Interaction (2026.findings-acl)

Copied to clipboard

Challenge: Current benchmarks evaluate task accuracy but overlook how agents interact . Preference-aware agents show 7.6% average UX improvement and 18.5% gain in preference alignment.
Approach: They propose a configurable environment that evaluates both what agents accomplish and how they interact.
Outcome: The proposed model improves performance and improves user experience by 7.6% and 18.5% respectively.
CiteEval: Principle-Driven Citation Evaluation for Source Attribution (2025.acl-long)

Copied to clipboard

Challenge: Current evaluation frameworks rely on NLI to assess binary or ternary support from cited sources, which is suboptimal for citation evaluation.
Approach: They propose a citation evaluation framework based on fine-grained citation ratings within a broad context and construct a multi-domain benchmark with high-quality human annotations.
Outcome: The proposed framework provides a high-quality human annotation benchmark and a suite of model-based metrics that exhibit strong correlation with human judgments.
Pub-LawBench: Public-Oriented Benchmarking for LegalAI (2026.acl-long)

Copied to clipboard

Challenge: Existing evaluation frameworks focus on legal professionals, not legal professionals.
Approach: They propose a public-oriented LegalAI benchmark grounded in legal functionalism and genre analysis to address this gap.
Outcome: The proposed model evaluates 17 large language models on Pub-LawBench using simple prompts and Chain-of-Thought under a vanilla inference setting.
CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches lack robustness to handle complex edge cases and generalizability across different domains.
Approach: They develop an accurate and lightweight verifier model for evaluation and outcome reward that matches unstructured outputs against standard answers.
Outcome: The proposed model can process multiple answer types including multi-subproblems, formulas, and sequence answers while identifying abnormal/invalid responses.
Reasoning Is Not All You Need: Examining LLMs for Multi-Turn Mental Health Conversations (2026.acl-long)

Copied to clipboard

Challenge: Existing evaluation frameworks focus on diagnostic accuracy and win-rates and often overlook alignment with patient-specific goals, values, and personalities required for meaningful conversations.
Approach: They propose a framework for synthetically generating realistic, multi-turn mental health sensemaking conversations and a dataset to examine their models in healthcare settings.
Outcome: The proposed framework synthesizes a dataset comprising over 2,200 patient–LLM conversations and evaluates them using human-centric criteria.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations